Papers with writing systems

3 papers
Robust Neural Machine Translation for Abugidas by Glyph Perturbation (2024.eacl-short)

Copied to clipboard

Challenge: Neural machine translation systems are vulnerable when trained on limited data.
Approach: They propose to add noise to the training phase to increase robustness of NMT systems trained on limited data.
Outcome: The proposed training strategy overcomes noise and improves robustness for low-resource tasks for abugida glyphs.
MiLiC-Eval: Benchmarking Multilingual LLMs for China’s Minority Languages (2025.findings-acl)

Copied to clipboard

Challenge: Large language models excel in high-resource languages but struggle with low-resourced languages . minority languages such as Tibetan, Uyghur, Kazakh, and Mongolian are marginalized in NLP research due to limited digital representation and the scarcity of training data.
Approach: They propose a benchmark for minority languages in China that tracks the progress of large language models on low-resource languages.
Outcome: The proposed benchmark focuses on underrepresented writing systems and syntax-intensive tasks.
Extensions to Brahmic script processing within the Nisaba library: new scripts, languages and utilities (2022.lrec-1)

Copied to clipboard

Challenge: a brahmic script is used to record endangered languages such as Dogri and Bengali for low-resource languages such that do not require visual normalization.
Approach: They propose to extend Brahmic script functionality within the Nisaba library of finite-state script normalization and processing utilities.
Outcome: The proposed extensions extend coverage from the original ten scripts to an additional ten of South Asia and beyond, including some used to record endangered languages such as Dogri.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations